<!DOCTYPE html>
<html class="client-nojs vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-0 vector-toc-not-available vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-0 skin-theme-clientpref-day vector-sticky-header-enabled" lang="de" dir="ltr"><head>
<meta charset="UTF-8">
<title>Clusteranalyse</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="icon" type="image/png" href="./_res_/favicon.png">
<link rel="canonical" href="https://de.wikipedia.org/wiki/Clusteranalyse"> <link href="./_mw_/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.math.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.wikimediamessages.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./_mw_/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/skins.vector.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link href="./_mw_/ext.gadget.citeRef.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.defaultPlainlinks.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiCommonHide.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiCommonLayout.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiCommonStyle.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiDarkmode.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiResponsive.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.specialSearch.css" rel="stylesheet" type="text/css">
<link rel="stylesheet" type="text/css" href="./_mw_/site.styles.css">
<link rel="stylesheet" type="text/css" href="./_mw_/noscript.css">
<link rel="stylesheet" type="text/css" href="./_res_/footer.css">
<link rel="stylesheet" type="text/css" href="./_res_/vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Clusteranalyse rootpage-Clusteranalyse skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading"><span class="mw-page-title-main">Clusteranalyse</span></h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="contentSub">
<div id="mw-content-subtitle"></div>
</div>
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="de" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="de" dir="ltr">
<p>Unter <b>Clusteranalyse</b> (<b>Clustering-Algorithmus</b>, gelegentlich auch: <b>Ballungsanalyse</b>) versteht man ein Verfahren zur Entdeckung von <i>Ähnlichkeitsstrukturen</i> in (meist relativ großen) Datenbeständen. Die so gefundenen Gruppen von „ähnlichen“ Objekten werden als <i><a href="Cluster_(Datenanalyse)" title="Cluster (Datenanalyse)">Cluster</a></i> bezeichnet, die Gruppenzuordnung als <i>Clustering.</i> Die gefundenen Ähnlichkeitsgruppen können graphentheoretisch, hierarchisch, partitionierend oder optimierend sein. Die Clusteranalyse ist eine wichtige Disziplin des <a href="Data-Mining" title="Data-Mining">Data-Minings</a>, des Analyseschritts des <a href="Knowledge_Discovery_in_Databases" title="Knowledge Discovery in Databases">Knowledge-Discovery-in-Databases-Prozesses</a>. Das Ziel der Clusteranalyse ist, <i>neue</i> Gruppen in den Daten zu identifizieren (im Gegensatz zur <a href="Klassifikation" title="Klassifikation">Klassifikation</a>, bei der Daten bestehenden Klassen zugeordnet werden). Man spricht von einem „uninformierten Verfahren“, da es nicht auf Klassen-Vorwissen angewiesen ist. Diese neuen Gruppen können anschließend beispielsweise zur automatisierten <a href="Klassifizierung" title="Klassifizierung">Klassifizierung</a>, zur <a href="Mustererkennung" title="Mustererkennung">Erkennung von Mustern</a> in der <a href="Bildverarbeitung" title="Bildverarbeitung">Bildverarbeitung</a> oder zur <a href="Marktsegmentierung" title="Marktsegmentierung">Marktsegmentierung</a> eingesetzt werden (oder in beliebigen anderen Verfahren, die auf ein derartiges Vorwissen angewiesen sind).
</p><p>Die zahlreichen <a href="Algorithmus" title="Algorithmus">Algorithmen</a> unterscheiden sich vor allem in ihrem Ähnlichkeits- und Gruppenbegriff, ihrem <a href="Cluster_(Datenanalyse)#Modelle_von_Clustern" title="Cluster (Datenanalyse)">Cluster-Modell</a>, ihrem algorithmischen Vorgehen (und damit ihrer Komplexität) und der Toleranz gegenüber Störungen in den Daten. Ob das von einem solchen Algorithmus generierte „Wissen“ nützlich ist, kann jedoch in der Regel nur ein Experte beurteilen. Ein Clustering-Algorithmus kann unter Umständen vorhandenes Wissen reproduzieren (beispielsweise Personendaten in die bekannten Gruppen „männlich“ und „weiblich“ unterteilen) oder auch für den Anwendungszweck nicht hilfreiche Gruppen generieren. Die gefundenen Gruppen lassen sich oft auch nicht verbal beschreiben (anders als z. B. bei „männliche Personen“), gemeinsame Eigenschaften werden in der Regel erst durch eine nachträgliche Analyse identifiziert. Bei der Anwendung von Clusteranalyse ist es daher oft notwendig, verschiedene Verfahren und verschiedene Parameter abzufragen, die Daten vorzuverarbeiten und beispielsweise Attribute auszuwählen oder wegzulassen.
</p>
<div class="mw-heading mw-heading2"><h2 id="Beispiel">Beispiel</h2></div>
<p>Angewendet auf einen Datensatz von Fahrzeugen könnte ein Clustering-Algorithmus (und eine nachträgliche Analyse der gefundenen Gruppen) beispielsweise folgende Struktur liefern:
</p>
<style data-mw-deduplicate="TemplateStyles:r249237068">
/* start https://de.wikipedia.org/ */
.mw-parser-output .stammbaum td{border-color:#000000!important}@media screen{html.skin-theme-clientpref-night .mw-parser-output .stammbaum td{border-color:#ffffff!important}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .stammbaum td{border-color:#ffffff!important}}
/* end https://de.wikipedia.org/ */
</style><table class="stammbaum" style="border-collapse:collapse; text-align:center;margin:1em;">
<tbody><tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Fahrzeuge</td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Fahrräder</td></tr>
<tr style="text-align:center;"></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2" style="border-bottom: 1px solid;"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;"><a href="LKW" class="mw-redirect" title="LKW">LKW</a></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;"><a href="PKW" class="mw-redirect" title="PKW">PKW</a></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;"><a href="Rikscha" title="Rikscha">Rikschas</a></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;"><a href="Rollermobil" title="Rollermobil">Rollermobile</a></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"></tr>
</tbody></table>
<p>Dabei ist Folgendes zu beachten:
</p>
<ul><li>Der Algorithmus selbst <i>liefert keine Interpretation</i> („LKW“) der gefundenen Gruppen. Hierzu ist eine separate Analyse der Gruppen notwendig.</li>
<li>Ein Mensch würde eine <a href="Fahrradrikscha" class="mw-redirect" title="Fahrradrikscha">Fahrradrikscha</a> als Untergruppe der Fahrräder ansehen. Für einen Clustering-Algorithmus aber sind 3 Räder oft ein signifikanter Unterschied, den sie mit einem Dreirad-<a href="Rollermobil" title="Rollermobil">Rollermobil</a> teilen.</li>
<li>Die Gruppen sind oft nicht „rein“, es können also beispielsweise kleine LKW in der Gruppe der PKW sein.</li>
<li>Es treten oft zusätzliche Gruppen auf, die nicht erwartet wurden („Polizeiautos“, „Cabrios“, „rote Autos“, „Autos mit Xenon-Scheinwerfern“).</li>
<li>Manche Gruppen werden nicht gefunden, beispielsweise „Motorräder“ oder „Liegeräder“.</li>
<li>Welche Gruppen gefunden werden, hängt <i>stark</i> vom verwendeten Algorithmus sowie von den Parametern und verwendeten Objekt-Attributen ab.</li>
<li>Oft wird auch nichts (Sinnvolles) gefunden.</li>
<li>In diesem Beispiel wurde nur <i>bekanntes</i> Wissen (wieder-)gefunden – als Verfahren zur „Wissensentdeckung“ ist die Clusteranalyse hier also gescheitert.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="Geschichte">Geschichte</h2></div>
<p>Historisch gesehen stammt das Verfahren aus der <a href="Taxonomie" title="Taxonomie">Taxonomie</a> in der <a href="Biologie" title="Biologie">Biologie</a>, wo über eine Clusterung von verwandten Arten eine Ordnung der Lebewesen ermittelt wird. Allerdings wurden dort ursprünglich keine automatischen Berechnungsverfahren eingesetzt. Inzwischen können zur Bestimmung der Verwandtschaft von Organismen unter anderem ihre <a href="Genom" title="Genom">Gensequenzen</a> verglichen werden (<i>Siehe auch:</i> <a href="Kladistik" title="Kladistik">Kladistik</a>).
Später wurde das Verfahren für die Zusammenhangsanalyse in die <a href="Sozialwissenschaften" title="Sozialwissenschaften">Sozialwissenschaften</a> eingeführt, weil es sich wegen des in den Gesellschaftswissenschaften in der Regel niedrigen <a href="Skalenniveau" title="Skalenniveau">Skalenniveaus</a> der Daten in diesen Disziplinen besonders eignet.<sup id="cite_ref-schlosser_1-0" class="reference"><a href="#cite_note-schlosser-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Prinzip/Funktionsweise"><span id="Prinzip.2FFunktionsweise"></span>Prinzip/Funktionsweise</h2></div>
<div class="mw-heading mw-heading3"><h3 id="Mathematische_Modellierung">Mathematische Modellierung</h3></div>
<p>Die Eigenschaften der zu untersuchenden Objekte werden mathematisch als <a href="Zufallsvariable" title="Zufallsvariable">Zufallsvariablen</a> aufgefasst. Sie werden in der Regel in Form von <a href="Vektor" title="Vektor">Vektoren</a> als Punkte in einem <a href="Vektorraum" title="Vektorraum">Vektorraum</a> dargestellt, deren Dimensionen die Eigenschaftsausprägungen des Objekts bilden. Bereiche, in denen sich Punkte anhäufen (<a href="Punktwolke" title="Punktwolke">Punktwolke</a>), werden Cluster genannt. Bei <a href="Streudiagramm" title="Streudiagramm">Streudiagrammen</a> dienen die Abstände der Punkte zueinander oder die Varianz innerhalb eines Clusters als sogenannte Proximitätsmaße, welche die Ähnlichkeit bzw. Unterschiedlichkeit zwischen den Objekten zum Ausdruck bringen.
</p><p>Ein Cluster kann auch als eine Gruppe von Objekten definiert werden, die <i>in Bezug auf einen berechneten <a href="Geometrischer_Schwerpunkt" title="Geometrischer Schwerpunkt">Schwerpunkt</a></i> eine minimale Abstandssumme haben. Dazu ist die Wahl eines <a href="Distanzma%C3%9F" class="mw-redirect" title="Distanzmaß">Distanzmaßes</a> erforderlich. In bestimmten Fällen sind die Abstände (bzw. umgekehrt die <a href="%C3%84hnlichkeitsma%C3%9F" class="mw-redirect" title="Ähnlichkeitsmaß">Ähnlichkeiten</a>) der Objekte untereinander direkt bekannt, sodass sie nicht aus der Darstellung im Vektorraum ermittelt werden müssen.
</p>
<div class="mw-heading mw-heading3"><h3 id="Grundsätzliche_Vorgehensweise"><span id="Grunds.C3.A4tzliche_Vorgehensweise"></span>Grundsätzliche Vorgehensweise</h3></div>
<p>In einer Gesamtmenge/-gruppe mit unterschiedlichen Objekten werden die Objekte, die sich ähnlich sind, zu Gruppen (Clustern) zusammengefasst.
</p><p>Als Beispiel sei folgende Musikanalyse gegeben. Die Werte ergeben sich aus dem Anteil der Musikstücke, die der Nutzer pro Monat online kauft.
</p>
<table class="wikitable">
<tbody><tr>
<th></th>
<th>Person 1</th>
<th>Person 2</th>
<th>Person 3
</th></tr>
<tr>
<td>Pop</td>
<td>2</td>
<td>1</td>
<td>8
</td></tr>
<tr>
<td>Rock</td>
<td>10</td>
<td>8</td>
<td>1
</td></tr>
<tr>
<td>Jazz</td>
<td>3</td>
<td>3</td>
<td>1
</td></tr></tbody></table>
<p>In diesem Beispiel würde man die Personen intuitiv in zwei Gruppen einteilen. Gruppe 1 besteht aus Person 1&2 und Gruppe 2 besteht aus Person 3. Das würden auch die meisten Clusteralgorithmen machen. Dieses Beispiel ist lediglich aufgrund der gewählten Werte so eindeutig, spätestens mit näher zusammenliegenden Werten und mehr Variablen (hier Musikrichtungen) und Objekten (hier Personen) ist leicht vorstellbar, dass die Einteilung in Gruppen nicht mehr so trivial ist.
</p><p>Etwas genauer und abstrakter ausgedrückt: Die Objekte einer <a href="Heterogenit%C3%A4t_(Naturwissenschaft)" class="mw-redirect" title="Heterogenität (Naturwissenschaft)">heterogenen</a> (beschrieben durch unterschiedliche Werte ihrer Variablen) Gesamtmenge werden mit Hilfe der Clusteranalyse zu Teilgruppen (Clustern/Segmenten) zusammengefasst, die in sich möglichst <a href="Homogenit%C3%A4t_(Physik)" class="mw-redirect" title="Homogenität (Physik)">homogen</a> (die Unterschiede der Variablen möglichst gering) sind. In unserem Musikbeispiel kann die Gruppe aller Musikhörer (eine sehr heterogene Gruppe) in die Gruppen der Jazzhörer, Rockhörer, Pophörer etc. (jeweils relativ homogen) unterteilt werden bzw. werden die Hörer mit ähnlichen Präferenzen zu der entsprechenden Gruppe zusammengefasst.
</p><p>Eine Clusteranalyse (z. B. bei der Marktsegmentierung) erfolgt dabei in folgenden Schritten:
</p><p>Schritte zur Clusterbildung:
</p>
<ol><li>Variablenauswahl: Auswahl (und Erhebung) der für die Untersuchung geeigneten Variablen. Sofern die Variablen der Objekte/Elemente noch nicht bekannt/vorgegeben sind, müssen alle für die Untersuchung wichtigen Variablen bestimmt und anschließend ermittelt werden.</li>
<li>Proximitätsbestimmung: Wahl eines geeigneten Proximitätsmaßes und Bestimmung der Distanz- bzw. Ähnlichkeitswerte (je nach Proximitätsmaß) zwischen den Objekten über das Proximitätsmaß. Abhängig von der Art der Variablen bzw. der Skalenart der Variablen wird eine entsprechende <a href="Distanzfunktion" title="Distanzfunktion">Distanzfunktion</a> zur Bestimmung des Abstandes (Distanz) zweier Elemente oder eine Ähnlichkeitsfunktion zur Bestimmung der Ähnlichkeit verwendet. Die Variablen werden zunächst einzeln verglichen und aus der Distanz der einzelnen Variablen die Gesamtdistanz (oder Ähnlichkeit) berechnet. Die Funktion zur Bestimmung von Distanz oder Ähnlichkeit wird auch Proximitätsmaß genannt. Der durch ein Proximitätsmaß ermittelte Distanz- bzw. Ähnlichkeitswert nennt sich Proximität. Werden alle Objekte miteinander verglichen, ergibt sich eine Proximitätsmatrix, die jeweils zwei Objekten eine Proximität zuweist.</li>
<li>Clusterbildung: Bestimmung und Durchführung eines/der geeigneten Clusterverfahren(s), um anschließend mit Hilfe des/dieser Verfahren(s) Gruppen/Cluster bilden zu können (die Proximitätsmatrix wird hierdurch reduziert). Im Regelfall werden hierbei mehrere Verfahren kombiniert, z. B.:
<ul><li>Finden von Ausreißern durch das Single-Linkage-Verfahren (hierarchisches Verfahren)</li>
<li>Bestimmung der Clusterzahl durch das Ward-Verfahren (hierarchisches Verfahren)</li>
<li>Bestimmung der Clusterzusammensetzung durch das Austauschverfahren (partitionierendes Verfahren)</li></ul></li></ol>
<p>Weitere Schritte der Clusteranalyse:
</p>
<ol><li>Bestimmung der Clusterzahl durch Betrachtung der Varianz innerhalb und zwischen den Gruppen. Hier wird bestimmt, wie viele Gruppen tatsächlich gebildet werden, denn bei der Clusterung selbst ist keine Abbruchbedingung vorgegeben. Z. B. wird bei einem agglomerativen Verfahren folglich so lange fusioniert, bis nur noch eine Gesamtgruppe vorhanden ist.</li>
<li>Interpretation der Cluster (abhängig von den inhaltlichen Ergebnissen, z. B. durch t-Wert)</li>
<li>Beurteilung der Güte der Clusterlösung (Bestimmung der Trennschärfe der Variablen, Gruppenstabilität)</li></ol>
<div class="mw-heading mw-heading3"><h3 id="Unterschiedliche_Proximitätsmaße/Skalen"><span id="Unterschiedliche_Proximit.C3.A4tsma.C3.9Fe.2FSkalen"></span>Unterschiedliche Proximitätsmaße/Skalen</h3></div>
<p>Die Skalen bezeichnen den Wertebereich, den die betrachtete(n) Variable(n) des Objektes annehmen kann/können. Je nachdem, welche Art von Skala vorliegt, muss man ein passendes Proximitätsmaß verwenden. Es gibt drei Hauptkategorien von Skalen:
</p>
<ul><li>Binäre Skalen<br>Die Variable kann zwei Werte annehmen, z. B. 0 oder 1, männlich oder weiblich.<br>Verwendete Proximitätsmaße (Beispiele):
<ul><li><a href="Jaccard-Koeffizient" title="Jaccard-Koeffizient">Jaccard-Koeffizient</a> (Ähnlichkeitsmaß)</li>
<li>Lance-Williams-Maß (Distanzmaß)</li></ul></li>
<li>Nominale Skalen<br>Die Variable kann unterschiedliche Werte annehmen, um eine qualitative Unterscheidung zu treffen, z. B., ob Pop, Rock oder Jazz bevorzugt wird.<br>Verwendete Proximitätsmaße (Beispiele):
<ul><li>Chi-Quadrat-Maß (Distanzmaß)</li>
<li><a href="Kontingenzkoeffizient" title="Kontingenzkoeffizient">Phi-Quadrat-Maß</a> (Distanzmaß)</li></ul></li>
<li>Metrische Skalen<br>Die Variable nimmt einen Wert auf einer vorher festgelegten Skala ein, um eine quantitative Aussage zu treffen, z. B., wie gerne eine Person Pop auf einer Skala von 1 bis 10 hört.<br>Verwendete Proximitätsmaße (Beispiele):
<ul><li>Pearson-<a href="Korrelationskoeffizient" class="mw-redirect" title="Korrelationskoeffizient">Korrelationskoeffizient</a> (Ähnlichkeitsmaß)</li>
<li><a href="Euklidische_Metrik" class="mw-redirect" title="Euklidische Metrik">Euklidische Metrik</a> (Distanzmaß)</li>
<li><a href="Minkowski-Metrik" class="mw-redirect" title="Minkowski-Metrik">Minkowski-Metrik</a> (Distanzmaß)</li></ul></li></ul>
<div class="mw-heading mw-heading3"><h3 id="Formen_der_Gruppenbildung_(Gruppenzugehörigkeit)"><span id="Formen_der_Gruppenbildung_.28Gruppenzugeh.C3.B6rigkeit.29"></span>Formen der Gruppenbildung (Gruppenzugehörigkeit)</h3></div>
<div class="hauptartikel" role="navigation"><span class="hauptartikel-pfeil" title="siehe" aria-hidden="true" role="presentation">→ </span><i><span class="hauptartikel-text">Hauptartikel</span>: <a href="Cluster_(Datenanalyse)" title="Cluster (Datenanalyse)">Cluster (Datenanalyse)</a></i></div>
<p>Es sind drei unterschiedliche Formen der Gruppenbildung (Gruppenzugehörigkeit) möglich. Bei den überlappenden Gruppen kann ein Objekt mehreren Gruppen zugeordnet werden, bei den nicht überlappenden Gruppen hingegen wird jedes Objekt nur einer einzelnen Gruppe (Segment, Cluster) zugeordnet. In den Fuzzy-Gruppen gehört ein Element jeder Gruppe mit einem bestimmten Grad des Zutreffens an.
</p>
<table class="stammbaum" style="border-collapse:collapse; text-align:center;margin:1em;">
<tbody><tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Nichtüberlappend</td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"></tr>
<tr style="font-weight:normal;"><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Formen der Gruppenbildung</td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Überlappend</td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Fuzzy</td></tr>
<tr style="text-align:center;"><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
</tbody></table>
<p>Man unterscheidet zwischen „harten“ und „weichen“ Clustering-Algorithmen. Harte Methoden (z. B. <a href="K-Means-Algorithmus" title="K-Means-Algorithmus">k-means</a>, <a href="Spektrales_Clustering" title="Spektrales Clustering">Spektrales Clustering</a>, Kernbasierte Hauptkomponentenanalyse (<i><span lang="en">kernel principal component analysis</span></i>, kurz: <i>kernel PCA</i>)) ordnen jeden Datenpunkt genau einem Cluster zu, wohingegen bei weichen Methoden (z. B. EM-Algorithmus mit <i><a href="Mischverteilung#Häufiger_Spezialfall:_Gaußsche_Mischmodelle" title="Mischverteilung">Gaußschen Mischmodellen</a></i> (<i><span lang="en">gaussian mixture models</span></i>, kurz: <i>GMMs</i>)) jedem Datenpunkt für jeden Cluster ein Grad zugeordnet wird, mit der dieser Datenpunkt in diesem Cluster zugeordnet werden kann. Weiche Methoden sind insbesondere dann nützlich, wenn die Datenpunkte relativ homogen im Raum verteilt sind und die Cluster nur als Regionen mit erhöhter Datenpunktdichte in Erscheinung treten, d. h., wenn es z. B. fließende Übergänge zwischen den Clustern oder Hintergrundrauschen gibt (harte Methoden sind in diesem Fall unbrauchbar).
</p>
<div class="mw-heading mw-heading3"><h3 id="Unterscheidung_der_Clusterverfahren">Unterscheidung der Clusterverfahren</h3></div>
<p>Clusterverfahren lassen sich in graphentheoretische, hierarchische, partitionierende und optimierende Verfahren sowie in weitere Unterverfahren einteilen.
</p>
<ul><li>Partitionierende Verfahren<br> verwenden eine gegebene Partitionierung und ordnen die Elemente durch Austauschfunktionen um, bis die verwendete Zielfunktion ein Optimum erreicht. Zusätzliche Gruppen können jedoch nicht gebildet werden, da die Anzahl der Cluster bereits am Anfang festgelegt wird (vgl. hierarchische Verfahren).</li>
<li>Hierarchische Verfahren<br> gehen von der feinsten (agglomerativ bzw. <i>bottom-up</i>) bzw. gröbsten (divisiv bzw. <i>top-down</i>) Partition aus (vgl. <a href="Top-down_und_Bottom-up" title="Top-down und Bottom-up">Top-down und Bottom-up</a>). Die gröbste Partition entspricht der Gesamtheit aller Elemente und die feinste Partition enthält lediglich ein Element bzw. jedes Element bildet seine eigene Gruppe/Partition. Durch Aufteilen bzw. Zusammenfassen lassen sich anschließend Cluster bilden. Einmal gebildete Gruppen können nicht mehr aufgelöst oder einzelne Elemente getauscht werden (vgl. partitionierende Verfahren). Agglomerative Verfahren kommen in der Praxis (z. B. bei der Marktsegmentierung im Marketing) sehr viel häufiger vor.</li></ul>
<table class="stammbaum" style="border-collapse:collapse; text-align:center;margin:1em;">
<tbody><tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">graphentheoretisch</td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">divisiv</td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">hierarchisch</td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">agglomerativ</td></tr>
<tr style="text-align:center;"><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Clusterverfahren</td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">Austauschverfahren</td></tr>
<tr style="text-align:center;"><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">partitionierend</td><td style="border-right: 1px solid; border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" style="border-right: 1px solid;"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2"><div style="width: 1em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">iterierte Minimaldistanzverfahren</td></tr>
<tr style="text-align:center;"><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="font-weight:normal;"><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-right: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td style="border-bottom: 1px solid;"><div style="width: 1em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td><td colspan="6" rowspan="2" style="border: 2px solid; padding: 0.2em; ;">optimierend</td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td><td rowspan="2" colspan="2"><div style="width: 2em; height: 2em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
<tr style="text-align:center;"><td colspan="2"><div style="width: 2em; height: 1em;"><span style="font: 1px/1px serif;"> </span></div></td></tr>
</tbody></table>
<p>Zu beachten ist, dass man noch diverse weitere Verfahren und Algorithmen unterscheiden kann, unter anderem überwachte (<i>supervised</i>) und nicht-überwachte (<i>unsupervised</i>) Algorithmen oder modellbasierte Algorithmen, bei denen eine Annahme über die zugrundeliegende Verteilung der Daten gemacht wird (z. B. <a href="GAUSSIAN" title="GAUSSIAN">Gaussian</a> mixture model).
</p>
<div class="mw-heading mw-heading2"><h2 id="Verfahren">Verfahren</h2></div>
<p>Es sind eine Vielzahl von Clustering-Verfahren in den unterschiedlichsten Anwendungsgebieten entwickelt worden. Man kann folgende Verfahrenstypen unterscheiden:
</p>
<ul><li>(zentrumsbasierte) partitionierende Verfahren,</li>
<li>hierarchische Verfahren,</li>
<li>dichte-basierte Verfahren,</li>
<li>gitter-basierte Verfahren und</li>
<li>kombinierte Verfahren.</li></ul>
<p>Die ersten beiden Verfahrenstypen sind die klassischen Clusterverfahren, während die anderen Verfahren eher neueren Datums sind.
</p>
<div class="mw-heading mw-heading3"><h3 id="Partitionierende_Clusterverfahren">Partitionierende Clusterverfahren</h3></div>
<style data-mw-deduplicate="TemplateStyles:r244148797">
/* start https://de.wikipedia.org/ */
.mw-parser-output .dewiki-gallery{text-align:center}.mw-parser-output .dewiki-gallery-title{font-size:110%;font-weight:bold}.mw-parser-output .dewiki-gallery-nav{font-size:110%}.mw-parser-output .dewiki-gallery-link{cursor:pointer;color:#3366cc}.mw-parser-output .dewiki-gallery-text{margin:0 .2em 0 .2em}@media(max-width:720px){.mw-parser-output .dewiki-gallery{float:none!important;clear:none!important;margin:0!important;display:flex;justify-content:center;flex-wrap:wrap;flex-direction:column}}
/* end https://de.wikipedia.org/ */
</style><div class="dewiki-gallery" style="border:1px solid transparent; float: right; clear: right; margin-left: 1em;"><div class="dewiki-gallery-title">Partitionierende Clusteranalyse</div><div class="dewiki-gallery-units"></div></div>
<p>Die Gemeinsamkeit der partitionierenden Verfahren ist, dass zunächst die Zahl der Cluster <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> festgelegt werden muss (Nachteil). Dann werden <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> Clusterzentren bestimmt und diese iterativ so lange verschoben, bis sich die Zuordnung der Beobachtungen zu den <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> Clusterzentren nicht mehr verändert, wobei eine vorgegebene <a href="Fehlerfunktion" title="Fehlerfunktion">Fehlerfunktion</a> minimiert wird. Ein Vorteil ist, dass Objekte während der Verschiebung der Clusterzentren ihre Clusterzugehörigkeit wechseln können.
</p>
<dl><dt><a href="K-Means-Algorithmus" title="K-Means-Algorithmus">k-Means-Algorithmus</a></dt>
<dd>Die Clusterzentren werden zufällig festgelegt und die Summe der quadrierten euklidischen Abstände der Objekte zu ihrem nächsten Clusterzentrum wird minimiert. Das Update der Clusterzentren geschieht durch Mittelwertbildung aller Objekte in einem Cluster.
<dl><dt><a href="K-Means-Algorithmus#K-Means++" title="K-Means-Algorithmus">k-Means++</a><sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup></dt>
<dd>Als Clusterzentren werden auch zufällig Objekte so ausgewählt, sodass sie etwa uniform im Raum der Objekte verteilt sind. Dies führt zu einem schnelleren Algorithmus.</dd>
<dt><a href="K-Means-Algorithmus#K-Median" title="K-Means-Algorithmus">k-Median-Algorithmus</a></dt>
<dd>Hier wird die Summe der <a href="Manhattan-Distanz" class="mw-redirect" title="Manhattan-Distanz">Manhattan-Distanzen</a> der Objekte zu ihrem nächsten Clusterzentrum minimiert. Das Update der Clusterzentren geschieht durch die Berechnung des Medians aller Objekte in einem Cluster. Ausreißer in den Daten haben dadurch weniger Einfluss.</dd>
<dt><a href="K-Means-Algorithmus#K-Medoids_(PAM)" title="K-Means-Algorithmus">k-Medoids oder Partitioning Around Medoids (PAM)</a><sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup></dt>
<dd>Die Clusterzentren sind hier immer Objekte. Durch Verschiebung von Clusterzentren auf ein benachbartes Objekt wird die Summe der Distanzen zum nächstgelegenen Clusterzentrum minimiert. Im Gegensatz zum k-Means-Verfahren werden nur die Distanzen zwischen den Objekten benötigt und nicht die Koordinaten der Objekte.</dd></dl></dd></dl>
<dl><dt><a href="Fuzzy-c-Means-Algorithmus" title="Fuzzy-c-Means-Algorithmus">Fuzzy-c-Means-Algorithmus</a><sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup></dt>
<dd>Für jedes Objekt wird ein Zugehörigkeitsgrad zu einem Cluster berechnet, oft aus dem reellwertigen Intervall [0,1] (Zugehörigkeitsgrad = 1: Objekt gehört vollständig zu einem Cluster, Zugehörigkeitsgrad = 0: Objekt gehört nicht zu dem Cluster). Dabei gilt: Je weiter ein Objekt vom Clusterzentrum entfernt ist, desto kleiner ist auch sein Zugehörigkeitsgrad zu diesem Cluster. Wie im k-Median-Verfahren werden die Clusterzentren dann verschoben, jedoch haben weit entfernte Objekte (kleiner Zugehörigkeitsgrad) einen geringeren Einfluss auf die Verschiebung als nahe Objekte. Damit wird auch eine <i>weiche</i> Clusterzuordnung erreicht: Jedes Objekt gehört zu jedem Cluster mit einem entsprechenden Zugehörigkeitsgrad.</dd>
<dt><a href="EM-Algorithmus#EM-Clustering" title="EM-Algorithmus">EM-Clustering</a><sup id="cite_ref-5" class="reference"><a href="#cite_note-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup></dt>
<dd>Die Cluster werden als <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> multivariate Normalverteilungen modelliert. Mit Hilfe des <a href="EM-Algorithmus" title="EM-Algorithmus">EM-Algorithmus</a> werden die unbekannten Parameter (<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \mu _{i},\Sigma _{i}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>μ<!-- μ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>,</mo>
<msub>
<mi mathvariant="normal">Σ<!-- Σ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \mu _{i},\Sigma _{i}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/bd108779ab10b7bf59bcc78ae77c73d52b679690.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:5.713ex; height:2.676ex;" alt="{\displaystyle \mu _{i},\Sigma _{i}}" loading="lazy"></span> mit <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle i=1,\ldots ,k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>i</mi>
<mo>=</mo>
<mn>1</mn>
<mo>,</mo>
<mo>…<!-- … --></mo>
<mo>,</mo>
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle i=1,\ldots ,k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/f2dd4cb0548150c5ac4440b8a1e3b4f6218bdfd1.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:11.453ex; height:2.509ex;" alt="{\displaystyle i=1,\ldots ,k}" loading="lazy"></span>) der Normalverteilungen iterativ geschätzt. Im Gegensatz zu k-Means wird damit eine <i>weiche</i> Clusterzuordnung erreicht: Mit einer gewissen Wahrscheinlichkeit gehört jedes Objekt zu jedem Cluster und jedes Objekt beeinflusst so die Parameter jedes Clusters.</dd>
<dt>Affinity-Propagation<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup></dt>
<dd>Affinity-Propagation (AP) ist ein deterministischer <a href="Nachrichtenaustausch" title="Nachrichtenaustausch">Message-Passing</a>-Algorithmus, der automatisch eine Anzahl Clusterzentren findet. Als Ähnlichkeitsfunktion kann je nach Art der Daten z. B. die negative euklidische Distanz verwendet werden. Abstände der Objekte zu ihrem nächsten Clusterzentrum werden minimiert. Das Update der Clusterzentren geschieht durch Berechnung der Responsibility und Availability aller Objekte gegeneinander sowie deren Summe, wobei ein zusätzlicher Dämpfungsfaktor numerische Instabilitäten vermeiden soll. Die Anzahl der Cluster wird automatisch ermittelt, das Ergebnis hängt aber von den manuell gesetzten Werten für Dämpfung und Selbstähnlichkeit der einzelnen Datenpunkte ab, sodass damit die ermittelte Anzahl der Cluster nicht die optimale sein muss.</dd></dl>
<div class="mw-heading mw-heading3"><h3 id="Hierarchische_Clusterverfahren">Hierarchische Clusterverfahren</h3></div>
<div class="hauptartikel" role="navigation"><span class="hauptartikel-pfeil" title="siehe" aria-hidden="true" role="presentation">→ </span><i><span class="hauptartikel-text">Hauptartikel</span>: <a href="Hierarchische_Clusteranalyse" title="Hierarchische Clusteranalyse">Hierarchische Clusteranalyse</a></i></div>
<div class="dewiki-gallery" style="border:1px solid transparent; float: right; clear: right; margin-left: 1em;"><div class="dewiki-gallery-title">Hierarchische Clusteranalyse</div><div class="dewiki-gallery-units"></div></div>
<p>Als hierarchische Clusteranalyse bezeichnet man eine bestimmte Familie von distanzbasierten Verfahren zur Clusteranalyse. Cluster bestehen hierbei aus Objekten, die zueinander eine geringere Distanz (oder umgekehrt: höhere Ähnlichkeit) aufweisen als zu den Objekten anderer Cluster. Dabei wird eine Hierarchie von Clustern aufgebaut: auf der einen Seite ein Cluster, der alle Objekte enthält, und auf der anderen Seite so viele Cluster, wie man Objekte hat, d. h., jedes Cluster enthält genau ein Objekt. Man unterscheidet zwei wichtige Typen von Verfahren:
</p>
<ul><li>Die divisiven Clusterverfahren, in denen zunächst alle Objekte als zu einem Cluster gehörig betrachtet werden und dann schrittweise die Cluster in immer kleinere Cluster aufgeteilt werden, bis jeder Cluster nur noch aus einem Objekt besteht (auch: „Top-down-Verfahren“)
<dl><dd><a href="Hierarchische_Clusteranalyse#Divisive_Berechnung" title="Hierarchische Clusteranalyse">Divisive Analysis Clustering</a> (DIANA)<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup>
<dl><dd>Man beginnt mit einem Cluster, das alle Objekte enthält. Im Cluster mit dem größten <i>Durchmesser</i> wird das Objekt gesucht, das die größte mittlere Distanz oder Unähnlichkeit zu den anderen Objekten des Clusters aufweist. Dies ist der Kern des Splitterclusters. Iterativ wird jedes Objekt, das nahe genug am Splittercluster ist, diesem hinzugefügt. Der gesamte Prozess wird wiederholt, bis jeder Cluster nur noch aus einem Objekt besteht.</dd></dl></dd></dl></li>
<li>Die agglomerativen Clusterverfahren, in denen zunächst jedes Objekt einen Cluster bildet und dann schrittweise die Cluster in immer größere Cluster zusammengefasst werden, bis alle Objekte zu einem Cluster gehören (auch: „Bottom-up-Verfahren“). Die Verfahren in dieser Familie unterscheiden zum einen nach den verwendeten <a href="%C3%84hnlichkeitsanalyse" title="Ähnlichkeitsanalyse">Distanz- bzw. Ähnlichkeitsmaßen</a> (zwischen Objekten, aber auch zwischen ganzen Clustern) und, meist wichtiger, nach ihrer <a href="Hierarchische_Clusteranalyse#Fusionierungsalgorithmen" title="Hierarchische Clusteranalyse">Fusionsvorschrift</a>, welche Cluster in einem Schritt zusammengefasst werden. Die Fusionierungsmethoden unterscheiden sich in der Art und Weise, wie die Distanz des fusionierten Clusters zu allen anderen Clustern berechnet wird. Wichtige Fusionierungsmethoden sind:
<dl><dd>Single Linkage<sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup>
<dl><dd>Die Cluster, deren nächste Objekte die kleinste Distanz oder Unähnlichkeit haben, werden fusioniert.</dd></dl></dd>
<dd>Ward Methode<sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup>
<dl><dd>Die Cluster, die den kleinsten Zuwachs der <a href="Totale_Varianz" title="Totale Varianz">totalen Varianz</a> haben, werden fusioniert.</dd></dl></dd></dl></li></ul>
<p>Für beide Verfahren gilt: einmal gebildete Cluster können nicht mehr verändert oder einzelne Objekte getauscht werden. Es wird die Struktur entweder stets nur verfeinert („divisiv“) oder nur verallgemeinert („agglomerativ“), sodass eine strikte Cluster-Hierarchie entsteht. An der entstandenen Hierarchie kann man nicht mehr erkennen, wie sie berechnet wurde.
</p>
<div class="mw-heading mw-heading3"><h3 id="Dichtebasierte_Verfahren">Dichtebasierte Verfahren</h3></div>
<div class="dewiki-gallery" style="border:1px solid transparent; float: right; clear: right; margin-left: 1em;"><div class="dewiki-gallery-title">Beispiele für dichtebasiertes Clustering</div><div class="dewiki-gallery-units"></div></div>
<p>Bei dichtebasiertem Clustering werden Cluster als Objekte in einem d-dimensionalen Raum betrachtet, die dicht beieinander liegen, getrennt durch Gebiete mit geringerer Dichte.
</p>
<dl><dt><a href="DBSCAN" title="DBSCAN">DBSCAN</a><sup id="cite_ref-13" class="reference"><a href="#cite_note-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup> (Density-Based Spatial Clustering of Applications with Noise)</dt>
<dd>Objekte, die in einem vorgegebenen Abstand <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \epsilon }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>ϵ<!-- ϵ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \epsilon }</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3837cad72483d97bcdde49c85d3b7b859fb3fd2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:0.944ex; height:1.676ex;" alt="{\displaystyle \epsilon }" loading="lazy"></span> mindestens <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> weitere Objekte haben, sind Kernobjekte. Zwei Kernobjekte, deren Distanz kleiner als <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \epsilon }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>ϵ<!-- ϵ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \epsilon }</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3837cad72483d97bcdde49c85d3b7b859fb3fd2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:0.944ex; height:1.676ex;" alt="{\displaystyle \epsilon }" loading="lazy"></span> ist, gehören dabei zum selben Cluster. Nicht-Kern-Objekte, die nahe einem Cluster liegen, werden diesem als Randobjekte hinzugefügt. Objekte, die weder Kernobjekte noch Randobjekte sind, sind Rauschobjekte.
<dl><dt><a href="OPTICS" title="OPTICS">OPTICS</a><sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup> (Ordering Points To Identify the Clustering Structure)</dt>
<dd>Der Algorithmus erweitert DBSCAN, sodass auch verschieden dichte Cluster erkannt werden. Die Wahl des Parameters <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \epsilon }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>ϵ<!-- ϵ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \epsilon }</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3837cad72483d97bcdde49c85d3b7b859fb3fd2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:0.944ex; height:1.676ex;" alt="{\displaystyle \epsilon }" loading="lazy"></span> ist nicht mehr so ausschlaggebend, um die Clusterstruktur der Objekte zu finden.</dd></dl></dd></dl>
<dl><dt>Maximum-Margin-Clustering<sup id="cite_ref-15" class="reference"><a href="#cite_note-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup></dt>
<dd>Es werden (leere) Bereiche im Raum der Objekte gesucht, die zwischen zwei Clustern liegen. Daraus werden Clustergrenzen bestimmt und damit auch die Cluster. Die Technik ist eng angebunden an <a href="Support_Vector_Machine" title="Support Vector Machine">Support-Vektor-Maschinen</a>.</dd></dl>
<div class="mw-heading mw-heading3"><h3 id="Gitterbasierte_Verfahren">Gitterbasierte Verfahren</h3></div>
<p>Bei gitterbasierten Clusterverfahren wird der Datenraum unabhängig von den Daten in endlich viele Zellen aufgeteilt. Der größte Vorteil dieses Ansatzes ist die geringe asymptotische Komplexität im Niedrigdimensionalen, da die Laufzeit von der Anzahl der Gitterzellen abhängt. Mit steigender Anzahl der Dimensionen wächst jedoch die Zahl der Gitterzellen exponentiell. Vertreter sind STING und CLIQUE. Zudem können Gitter zur Beschleunigung anderer Algorithmen eingesetzt werden, bspw. zur Approximation von k-means oder zur Berechnung von DBSCAN (GriDBSCAN).
</p>
<ul><li>Der Algorithmus <b>STING</b><sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> (STatistical INformation Grid-based Clustering) teilt den Datenraum rekursiv in rechteckige Zellen. Statistische Informationen für jede Zelle werden auf der untersten Rekursionsebene vorausberechnet. Relevante Zellen werden anschließend mit einem Top-Down Ansatz berechnet und zurückgegeben.</li>
<li><b>CLIQUE</b><sup id="cite_ref-17" class="reference"><a href="#cite_note-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> (CLustering In QUEst) arbeitet in zwei Phasen: Zunächst wird der Datenraum in dicht besetzte d-dimensionale Zellen partitioniert. Die zweite Phase bestimmt aus diesen Zellen größere Cluster, indem durch eine Greedy-Strategie größtmögliche Regionen dicht besetzter Zellen ermittelt werden.</li></ul>
<div class="mw-heading mw-heading3"><h3 id="Kombinierte_Verfahren">Kombinierte Verfahren</h3></div>
<p>In der Praxis werden oft auch Kombinationen von Verfahren benutzt. Ein Beispiel ist es erst eine hierarchische Clusteranalyse durchzuführen, um eine geeignete Clusterzahl zu bestimmen, und danach noch ein k-Means Clustering, um das Resultat des Clusterings zu verbessern. Oft lässt sich in speziellen Situation zusätzliche Information ausnutzen, sodass z. B. die Dimension oder die Anzahl der zu clusternden Objekte reduziert wird.
</p>
<dl><dt><a href="Spektrales_Clustering" title="Spektrales Clustering">Spektrales Clustering</a><sup id="cite_ref-18" class="reference"><a href="#cite_note-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup></dt>
<dd>Die zu clusterenden Objekte können auch als Knoten eines <a href="Graph_(Graphentheorie)" title="Graph (Graphentheorie)">Graphs</a> aufgefasst werden und die gewichteten Kanten geben Distanz oder Unähnlichkeit wieder. Die <a href="Laplace-Matrix" title="Laplace-Matrix">Laplace-Matrix</a>, eine spezielle Transformierte der <a href="Adjazenzmatrix" title="Adjazenzmatrix">Adjazenzmatrix</a> (Matrix der Ähnlichkeit zwischen allen <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle n}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>n</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle n}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/a601995d55609f2d9f5e233e36fbe9ea26011b3b.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.395ex; height:1.676ex;" alt="{\displaystyle n}" loading="lazy"></span> Objekten), hat bei <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> Zusammenhangskomponenten (Clustern) den <a href="Eigenwert" class="mw-redirect" title="Eigenwert">Eigenwert</a> Null mit der Vielfachheit <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span>. Daher untersucht man die <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> kleinsten Eigenwerte der Laplace-Matrix und den zugehörigen <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span>-dimensionalen Eigenraum (<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k\ll n}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
<mo>≪<!-- ≪ --></mo>
<mi>n</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k\ll n}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/7a25760f79269b8f0fb75a078a66bf88403f74bf.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:6.22ex; height:2.176ex;" alt="{\displaystyle k\ll n}" loading="lazy"></span>). Statt in einem hochdimensionalen Raum wird nun in dem niedrigdimensionalen Eigenraum, z. B. mit dem k-Means-Verfahren, geclustert.</dd>
<dt>Multiview-Clustering<sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup></dt>
<dd>Hierbei wird davon ausgegangen, dass man mehrere Distanz- oder Ähnlichkeitsmatrizen (sog. <i>views</i>) der Objekte generieren kann. Ein Beispiel sind Webseiten als zu clusternde Objekte: eine Distanzmatrix kann auf Basis der gemeinsam verwendeten Worte berechnet werden, eine zweite Distanzmatrix auf Basis der Verlinkung. Dann wird ein Clustering (oder ein Clustering-Schritt) mit der einen Distanzmatrix durchgeführt und das Ergebnis als Input für ein Clustering (oder ein Clustering-Schritt) mit der anderen Distanzmatrix benutzt. Dies wird wiederholt, bis sich die Clusterzugehörigkeit der Objekte stabilisiert.</dd>
<dt><a href="Balanced_iterative_reducing_and_clustering_using_hierarchies" class="mw-redirect" title="Balanced iterative reducing and clustering using hierarchies">Balanced iterative reducing and clustering using hierarchies</a> (BIRCH)</dt>
<dd>Für (sehr) große Datensätze wird zunächst ein Preclustering durchgeführt. Die so gewonnenen Cluster (nicht mehr die Objekte) werden dann z. B. mit einer hierarchischen Clusteranalyse weitergeclustert. Dies ist die Basis des eigens für SPSS entwickelten und dort eingesetzten <i>Two-Step clusterings</i>.</dd></dl>
<div class="mw-heading mw-heading3"><h3 id="Biclustering">Biclustering</h3></div>
<div class="hauptartikel" role="navigation"><span class="hauptartikel-pfeil" title="siehe" aria-hidden="true" role="presentation">→ </span><i><span class="hauptartikel-text">Hauptartikel</span>: <a href="Biclustering" title="Biclustering">Biclustering</a></i></div>
<p>Biclustering, Co-Clustering oder Two-Mode Clustering ist eine Technik, die das gleichzeitige Clustering von Zeilen und Spalten einer Matrix ermöglicht. Zahlreiche Biclustering-Algorithmen wurden für die Bioinformatik entwickelt, darunter: Block Clustering, CTWC (Coupled Two-Way Clustering), ITWC (Interrelated Two-Way Clustering), δ-Bicluster, δ-pCluster, δ-Pattern, FLOC, OPC, Plaid Model, OPSMs (Order-preserving Submatrices), Gibbs, SAMBA (Statistical-Algorithmic Method for Bicluster Analysis), Robust Biclustering Algorithm (RoBA), Crossing Minimization, cMonkey, PRMs, DCC, LEB (Localize and Extract Biclusters), QUBIC (QUalitative BIClustering), BCCA (Bi-Correlation Clustering Algorithm) und FABIA (Factor Analysis for Bicluster Acquisition). Auch in anderen Anwendungsgebieten werden Biclustering-Algorithmen vorgeschlagen und eingesetzt. Dort sind sie unter den Bezeichnungen Co-Clustering, Bidimensional Clustering sowie Subspace Clustering zu finden.
</p>
<div class="mw-heading mw-heading2"><h2 id="Evaluation">Evaluation</h2></div>
<p>Die Ergebnisse eines jeden Cluster-Verfahrens müssen wie bei allen Verfahren des maschinellen Lernens einer Evaluation unterzogen werden. Die Güte einer fertig berechneten Einteilung der Datenpunkte in Gruppen kann anhand verschiedener Metriken eingeschätzt werden. Die Auswahl einer geeigneten Metrik muss immer im Hinblick auf die verwendeten Daten, die vorgegebene Fragestellung sowie die gewählte Clustering-Methode erfolgen.
</p><p>Es wird zwischen internen, externen und relativen Metriken unterschieden.<sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup>
Interne Metriken bewerten die Cluster rein anhand des vorliegenden Datensatzes und dem internen Zusammenhang der Daten. Externe Metriken ziehen zusätzlich externe Daten (z. B. Experten Label) hinzu. Relative Metriken vergleichen die Ergebnisse zweier Cluster-Algorithmen.
</p>
<div class="mw-heading mw-heading3"><h3 id="Interne_Metriken">Interne Metriken</h3></div>
<p>Internen Validitätsmetriken haben den Vorteil, dass kein Wissen über die korrekte Zuordnung vorhanden sein muss, um das Ergebnis zu bewerten.
</p>
<div class="mw-heading mw-heading4"><h4 id="Silhouette_Score">Silhouette Score</h4></div>
<p>Zu den in der Praxis häufig verwendeten Evaluations-Kennzahlen gehört der 1987 von Peter J. Rousseeuw vorgestellte <a href="Silhouettenkoeffizient" title="Silhouettenkoeffizient">Silhouettenkoeffizient</a>. Dieser berechnet pro Datenpunkt einen Wert, der angibt wie gut die Zuordnung zur gewählten Gruppe im Vergleich zu allen deren Gruppen erfolgt ist.<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup> Mit einer Laufzeitkomplexität von <a href="Landau-Symbole" title="Landau-Symbole">O(n²)</a> ist der Silhouettenkoeffizient aber langsamer als viele Clusterverfahren, und kann so nicht auf großen Daten verwendet werden.
</p>
<div class="mw-heading mw-heading3"><h3 id="Externe_Metriken">Externe Metriken</h3></div>
<div class="mw-heading mw-heading4"><h4 id="Rand-Index">Rand-Index</h4></div>
<p>Der Rand-Index kann zur Bewertung der Übereinstimmung gefundener Cluster mit einer Ground-Truth herangezogen werden (z. B. können Clusterzugehörtigkeiten explizit in der <a href="Laplace-Matrix" title="Laplace-Matrix">Laplace-Matrix</a> oder der <a href="Adjazenzmatrix" title="Adjazenzmatrix">Adjazenzmatrix</a> festgehalten werden). Beim Rand-Index handelt es sich um die <a href="Korrektklassifikationsrate" class="mw-redirect" title="Korrektklassifikationsrate">Korrektklassifikationsrate</a> für die Klassifikation „das Paar A und B tritt gemeinsam im Cluster auf“.
</p>
<div class="mw-heading mw-heading4"><h4 id="Matthews_Korrelationskoeffizient">Matthews Korrelationskoeffizient</h4></div>
<div class="hauptartikel" role="navigation"><span class="hauptartikel-pfeil" title="siehe" aria-hidden="true" role="presentation">→ </span><i><span class="hauptartikel-text">Hauptartikel</span>: <a href="Matthews_Korrelationskoeffizient" class="mw-redirect" title="Matthews Korrelationskoeffizient">Matthews Korrelationskoeffizient</a></i></div>
<div class="mw-heading mw-heading4"><h4 id="Mutual_Information">Mutual Information</h4></div>
<div class="hauptartikel" role="navigation"><span class="hauptartikel-pfeil" title="siehe" aria-hidden="true" role="presentation">→ </span><i><span class="hauptartikel-text">Hauptartikel</span>: <a href="Mutual_Information" class="mw-redirect" title="Mutual Information">Mutual Information</a></i></div>
<div class="mw-heading mw-heading2"><h2 id="Siehe_auch">Siehe auch</h2></div>
<ul><li><a href="Faktorenanalyse" title="Faktorenanalyse">Faktorenanalyse</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="Literatur">Literatur</h2></div>
<div class="mw-heading mw-heading3"><h3 id="Grundlagen_und_Verfahren">Grundlagen und Verfahren</h3></div>
<ul><li>Martin Ester, Jörg Sander: <i>Knowledge Discovery in Databases. Techniken und Anwendungen.</i> Springer, Berlin 2000, ISBN 3-540-67328-8.</li>
<li>K. Backhaus, B. Erichson, W. Plinke, R. Weiber: <i>Multivariate Analysemethoden. Eine anwendungsorientierte Einführung.</i> Springer, 2003, ISBN 3-540-00491-2.</li>
<li>S. Bickel, T. Scheffer: <i>Multi-View Clustering.</i> In: <i>Proceedings of the IEEE International Conference on Data Mining.</i> 2004.</li>
<li>J. Shi, J. Malik: <i>Normalized Cuts and Image Segmentation.</i> In: <i>Proc. of IEEE Conf. on Comp. Vision and Pattern Recognition.</i> Puerto Rico 1997.</li>
<li><a href="Otto_Schlosser_(Sozialwissenschaftler)" title="Otto Schlosser (Sozialwissenschaftler)">Otto Schlosser</a>: <cite style="font-style:italic">Einführung in die sozialwissenschaftliche Zusammenhangsanalyse</cite>. Rowohlt, Reinbek bei Hamburg 1976, ISBN 3-499-21089-4.<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.au=Otto+Schlosser&rft.btitle=Einf%C3%BChrung+in+die+sozialwissenschaftliche+Zusammenhangsanalyse&rft.date=1976&rft.genre=book&rft.isbn=3499210894&rft.place=Reinbek+bei+Hamburg&rft.pub=Rowohlt" style="display:none"> </span></li>
<li>C. Hennig, M. Meila, F. Murtagh, R. Rocci: <i>Handbook of Cluster Analysis</i>. Chapman and Hall/CRC, ISBN 978-0-429-18547-2.</li></ul>
<div class="mw-heading mw-heading3"><h3 id="Anwendung">Anwendung</h3></div>
<ul><li>J. Bacher, A. Pöge, K. Wenzig: <i>Clusteranalyse – Anwendungsorientierte Einführung in Klassifikationsverfahren.</i> 3. Auflage. Oldenbourg, München 2010, ISBN 978-3-486-58457-8.</li>
<li>J. Bortz: <i>Statistik für Sozialwissenschaftler.</i> Springer, Berlin 1999, Kap. 16 Clusteranalyse</li>
<li>W. Härdle, L. Simar: <i>Applied Multivariate Statistical Analysis.</i> Springer, New York 2003.</li>
<li>C. Homburg, H. Krohmer: <i>Marketingmanagement: Strategie – Instrumente – Umsetzung – Unternehmensführung.</i> 3. Auflage. Gabler, Wiesbaden 2009, Kapitel 8.2.2</li>
<li>H. Moosbrugger, D. Frank: <i>Clusteranalytische Methoden in der Persönlichkeitsforschung. Eine anwendungsorientierte Einführung in taxometrische Klassifikationsverfahren.</i> Huber, Bern 1992, ISBN 3-456-82320-7.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="Einzelnachweise">Einzelnachweise</h2></div>
<ol class="references">
<li id="cite_note-schlosser-1"><span class="mw-cite-backlink"><a href="#cite_ref-schlosser_1-0">↑</a></span> <span class="reference-text">Vgl. unter anderem <a href="Otto_Schlosser_(Sozialwissenschaftler)" title="Otto Schlosser (Sozialwissenschaftler)">Otto Schlosser</a>: <cite style="font-style:italic">Einführung in die sozialwissenschaftliche Zusammenhangsanalyse</cite>. Rowohlt, Reinbek bei Hamburg 1976, ISBN 3-499-21089-4.<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.au=Otto+Schlosser&rft.btitle=Einf%C3%BChrung+in+die+sozialwissenschaftliche+Zusammenhangsanalyse&rft.date=1976&rft.genre=book&rft.isbn=3499210894&rft.place=Reinbek+bei+Hamburg&rft.pub=Rowohlt" style="display:none"> </span></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><a href="#cite_ref-2">↑</a></span> <span class="reference-text">D. Arthur, S. Vassilvitskii: <i>k-means++: The advantages of careful seeding.</i> In: <i>Proceedings of the eighteenth annual ACM-SIAM Symposium on Discrete algorithms.</i> Society for Industrial and Applied Mathematics, 2007, S. 1027–1035.</span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><a href="#cite_ref-3">↑</a></span> <span class="reference-text">S. Vinod: <cite style="font-style:italic">Integer programming and the theory of grouping</cite>. In: <cite style="font-style:italic">Journal of the American Statistical Association</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em"> </span>64</span>, 1969, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em"> </span>506–517</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1080/01621459.1969.10500990">10.1080/01621459.1969.10500990</a></span>, <a href="JSTOR" title="JSTOR">JSTOR</a>:<a rel="nofollow" class="external text" href="http://www.jstor.org/stable/2283635">2283635</a>.<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.atitle=Integer+programming+and+the+theory+of+grouping&rft.au=S.+Vinod&rft.btitle=Journal+of+the+American+Statistical+Association&rft.date=1969&rft.doi=10.1080%2F01621459.1969.10500990&rft.genre=book&rft.pages=506-517&rft.volume=64" style="display:none"> </span></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><a href="#cite_ref-4">↑</a></span> <span class="reference-text">J.C. Bezdek: <cite style="font-style:italic">Pattern recognition with fuzzy objective function algorithms</cite>. Plenum Press, New York 1981.<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.au=J.C.+Bezdek&rft.btitle=Pattern+recognition+with+fuzzy+objective+function+algorithms&rft.date=1981&rft.genre=book&rft.place=New+York&rft.pub=Plenum+Press" style="display:none"> </span></span>
</li>
<li id="cite_note-5"><span class="mw-cite-backlink"><a href="#cite_ref-5">↑</a></span> <span class="reference-text">A. P. Dempster, N. M. Laird, D. B. Rubin: <i>Maximum Likelihood from Incomplete Data via the EM algorithm.</i> In: <i>Journal of the Royal Statistical Society.</i> Series B, 39(1), 1977, S. 1–38, <a href="https://doi.org/10.1111/j.2517-6161.1977.tb01600.x" class="extiw external" title="doi:10.1111/j.2517-6161.1977.tb01600.x">doi:10.1111/j.2517-6161.1977.tb01600.x</a>.</span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><a href="#cite_ref-6">↑</a></span> <span class="reference-text">B. J. Frey, D. Dueck: <i>Clustering by passing messages between data points.</i> In: <i><a href="Science" title="Science">Science</a>.</i> Band 315, Nr. 5814, 2007, S. 972–976, <a href="https://doi.org/10.1126/science.1136800" class="extiw external" title="doi:10.1126/science.1136800">doi:10.1126/science.1136800</a>.</span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><a href="#cite_ref-7">↑</a></span> <span class="reference-text">L. Kaufman, P. J. Roussew: <i>Finding Groups in Data – An Introduction to Cluster Analysis.</i> John Wiley & Sons, 1990.</span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><a href="#cite_ref-8">↑</a></span> <span class="reference-text">K. Florek, J. Łukasiewicz, J. Perkal, H. Steinhaus, S. Zubrzycki: <i>Taksonomia wrocławska.</i> In: <i>Przegląd Antropol.</i> Band 17, 1951, S. 193–211.</span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><a href="#cite_ref-9">↑</a></span> <span class="reference-text">K. Florek, J. Łukaszewicz, J. Perkal, H. Steinhaus, S. Zubrzycki: <i>Sur la liaison et la division des points d'un ensemble fini.</i> In: <i>Colloquium Mathematicae.</i> Vol. 2, No. 3–4, Institute of Mathematics Polish Academy of Sciences 1951, S. 282–285.</span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><a href="#cite_ref-10">↑</a></span> <span class="reference-text">L. L. McQuitty: <i>Elementary linkage analysis for isolating orthogonal and oblique types and typal relevancies.</i> In: <i>Educational and Psychological Measurement.</i> 1957, S. 207–229, <a href="https://doi.org/10.1177/001316445701700204" class="extiw external" title="doi:10.1177/001316445701700204">doi:10.1177/001316445701700204</a>.</span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><a href="#cite_ref-11">↑</a></span> <span class="reference-text">P. H. Sneath: <i>The application of computers to taxonomy.</i> In: <i>Journal of General Microbiology.</i> Band 17, Nr. 1, 1957, S. 201–226, <a href="https://doi.org/10.1099/00221287-17-1-201" class="extiw external" title="doi:10.1099/00221287-17-1-201">doi:10.1099/00221287-17-1-201</a>.</span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><a href="#cite_ref-12">↑</a></span> <span class="reference-text">J. H. Ward Jr.: <i>Hierarchical grouping to optimize an objective function.</i> In: <i>Journal of the American Statistical Association.</i> Band 58, Nr. 301, 1963, S. 236–244, <a href="https://doi.org/10.1080/01621459.1963.10500845" class="extiw external" title="doi:10.1080/01621459.1963.10500845">doi:10.1080/01621459.1963.10500845</a>, <a href="JSTOR" title="JSTOR">JSTOR</a>:<a rel="nofollow" class="external text" href="http://www.jstor.org/stable/2282967">2282967</a>.</span>
</li>
<li id="cite_note-13"><span class="mw-cite-backlink"><a href="#cite_ref-13">↑</a></span> <span class="reference-text">M. Ester, H. P. Kriegel, J. Sander, X. Xu: <i>A density-based algorithm for discovering clusters in large spatial databases with noise.</i> In: <i>KDD-96 Proceedings.</i> Vol. 96, 1996, S. 226–231.</span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><a href="#cite_ref-14">↑</a></span> <span class="reference-text">Mihael Ankerst, Markus M. Breunig, <a href="Hans-Peter_Kriegel" title="Hans-Peter Kriegel">Hans-Peter Kriegel</a>, Jörg Sander: <cite style="font-style:italic">OPTICS: Ordering Points To Identify the Clustering Structure</cite>. In: <cite style="font-style:italic">ACM SIGMOD international conference on Management of data</cite>. ACM Press, 1999, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em"> </span>49–60</span> (<a rel="nofollow" class="external text" href="https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.129.6542">CiteSeerX</a> [abgerufen am 3. Februar 2023]).<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.atitle=OPTICS%3A+Ordering+Points+To+Identify+the+Clustering+Structure&rft.au=Mihael+Ankerst%2C+Markus+M.+Breunig%2C+Hans-Peter+Kriegel%2C+...&rft.btitle=ACM+SIGMOD+international+conference+on+Management+of+data&rft.date=1999&rft.genre=book&rft.pages=49-60&rft.pub=ACM+Press" style="display:none"> </span></span>
</li>
<li id="cite_note-15"><span class="mw-cite-backlink"><a href="#cite_ref-15">↑</a></span> <span class="reference-text">L. Xu, J. Neufeld, B. Larson, D. Schuurmans: <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper_files/paper/2004/file/6403675579f6114559c90de0014cd3d6-Paper.pdf"><i>Maximum margin clustering.</i></a> In: <i>Advances in neural information processing systems.</i> 2004, S. 1537–1544.</span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><a href="#cite_ref-16">↑</a></span> <span class="reference-text"><span class="cite">Wei Wang, Jiog Yang, Richard R. Muntz: <a rel="nofollow" class="external text" href="https://dl.acm.org/doi/abs/10.5555/645923.758369"><i>STING: A Statistical Information Grid Approach to Spatial Data Mining.</i></a> In: <i>Proceedings of the 23rd International Conference on Very Large Data Bases.</i> 25. August 1997,<span class="Abrufdatum"> abgerufen am 3. Februar 2023</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&rfr_id=info%3Asid%2Fde.wikipedia.org%3AClusteranalyse&rft.title=STING%3A+A+Statistical+Information+Grid+Approach+to+Spatial+Data+Mining&rft.description=STING%3A+A+Statistical+Information+Grid+Approach+to+Spatial+Data+Mining&rft.identifier=https%3A%2F%2Fdl.acm.org%2Fdoi%2Fabs%2F10.5555%2F645923.758369&rft.creator=Wei+Wang%2C+Jiog+Yang%2C+Richard+R.+Muntz&rft.date=1997-08-25&rft.language=en"> </span></span>
</li>
<li id="cite_note-17"><span class="mw-cite-backlink"><a href="#cite_ref-17">↑</a></span> <span class="reference-text"><span class="cite">Rakesh Agrawal, Johannes Gehrke, Dimitrios Gunopulos, Prabhakar Raghavan: <a rel="nofollow" class="external text" href="https://dl.acm.org/doi/abs/10.1145/276304.276314"><i>Automatic subspace clustering of high dimensional data for data mining applications.</i></a> In: <i>Proceedings of the 1998 ACM SIGMOD international conference on Management of data.</i><span class="Abrufdatum"> Abgerufen am 3. Februar 2023</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&rfr_id=info%3Asid%2Fde.wikipedia.org%3AClusteranalyse&rft.title=Automatic+subspace+clustering+of+high+dimensional+data+for+data+mining+applications&rft.description=Automatic+subspace+clustering+of+high+dimensional+data+for+data+mining+applications&rft.identifier=https%3A%2F%2Fdl.acm.org%2Fdoi%2Fabs%2F10.1145%2F276304.276314&rft.creator=Rakesh+Agrawal%2C+Johannes+Gehrke%2C+Dimitrios+Gunopulos%2C+Prabhakar+Raghavan&rft.language=en"> </span></span>
</li>
<li id="cite_note-18"><span class="mw-cite-backlink"><a href="#cite_ref-18">↑</a></span> <span class="reference-text">W. E. Donath, A. J. Hoffman: <i>Lower bounds for the partitioning of graphs.</i> In: <i>IBM Journal of Research and Development.</i> Band 17, Nr. 5, 1973, S. 420–425, <a href="https://doi.org/10.1147/rd.175.0420" class="extiw external" title="doi:10.1147/rd.175.0420">doi:10.1147/rd.175.0420</a>.</span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><a href="#cite_ref-19">↑</a></span> <span class="reference-text">M. Fiedler: <i>Algebraic connectivity of graphs.</i> In: <i>Czechoslovak Mathematical Journal.</i> Band 23, Nr. 2, 1973, S. 298–305.</span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><a href="#cite_ref-20">↑</a></span> <span class="reference-text">S. Bickel, T. Scheffer: <i>Multi-View Clustering.</i> In: <i>ICDM.</i> Vol. 4, Nov 2004, S. 19–26, <a href="https://doi.org/10.1109/ICDM.2004.10095" class="extiw external" title="doi:10.1109/ICDM.2004.10095">doi:10.1109/ICDM.2004.10095</a>.</span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><a href="#cite_ref-21">↑</a></span> <span class="reference-text">Mohammed J. Zaki, Wagner Meira Jr.: <cite style="font-style:italic">Data mining and analysis: fundamental concepts and algorithms</cite>. Cambridge University Press, 2014, ISBN 978-0-511-81011-4, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em"> </span>425<span style="display:inline-block;width:.2em"> </span>ff</span>. (<a rel="nofollow" class="external text" href="https://doc.lagout.org/Others/Data%20Mining/Data%20Mining%20and%20Analysis_%20Fundamental%20Concepts%20and%20Algorithms%20%5BZaki%20%26%20Meira%202014-05-12%5D.pdf">lagout.org</a> [PDF; abgerufen am 3. Februar 2023]).<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.au=Mohammed+J.+Zaki%2C+Wagner+Meira+Jr.&rft.btitle=Data+mining+and+analysis%3A+fundamental+concepts+and+algorithms&rft.date=2014&rft.genre=book&rft.isbn=9780511810114&rft.pages=425+ff.&rft.pub=Cambridge+University+Press" style="display:none"> </span></span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><a href="#cite_ref-22">↑</a></span> <span class="reference-text">Peter J.Rousseeuw: <cite style="font-style:italic">Silhouettes: A graphical aid to the interpretation and validation of cluster analysis</cite>. In: <cite style="font-style:italic">Journal of Computational and Applied Mathematics</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em"> </span>20</span>. Elsevier, November 1987, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em"> </span>53–65</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1016/0377-0427%2887%2990125-7">10.1016/0377-0427(87)90125-7</a></span>.<span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&rfr_id=info:sid/de.wikipedia.org:Clusteranalyse&rft.atitle=Silhouettes%3A+A+graphical+aid+to+the+interpretation+and+validation+of+cluster+analysis&rft.au=Peter+J.Rousseeuw&rft.btitle=Journal+of+Computational+and+Applied+Mathematics&rft.date=1987-11&rft.doi=10.1016%2F0377-0427%2887%2990125-7&rft.genre=book&rft.pages=53-65&rft.pub=Elsevier&rft.volume=20" style="display:none"> </span></span>
</li>
</ol>
<div class="hintergrundfarbe1 rahmenfarbe1 navigation-not-searchable normdaten-typ-s" style="border-style: solid; border-width: 1px; clear: left; margin-bottom:1em; margin-top:1em; padding: 0.25em; overflow: hidden; word-break: break-word; word-wrap: break-word;" id="normdaten">
<div style="display: table-cell; vertical-align: middle; width: 100%;">
<div>
Normdaten (Sachbegriff): <a href="Gemeinsame_Normdatei" title="Gemeinsame Normdatei">GND</a>: <span class="-print"><a rel="nofollow" class="external text" href="https://d-nb.info/gnd/4070044-6">4070044-6</a></span> </div>
</div></div></div><!--htdig_noindex--><div><div class="zim-footer">
Dieser Artikel wurde von <a class="external text" title="Zuletzt bearbeitet am 2025-12-24" href="https://de.wikipedia.org/wiki/?title=Clusteranalyse&oldid=262706180">Wikipedia</a> herausgegeben. Der Text ist unter <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.de">Creative Commons Attribution-Share Alike 4.0</a> verfügbar, sofern nicht anders angegeben. Für die Mediendateien können zusätzliche Bedingungen gelten.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
<script src="./_webp_/webpHandler.js"></script>
</body></html>